Papers with aggregated first-order influence approximation
AutoMixer: Checkpoint Artifacts as Automatic Data Mixers (2025.acl-long)
Copied to clipboard
| Challenge: | In language model training, it is difficult to obtain the right data mixtures for various tasks as the relationship between data and tasks is difficult. |
| Approach: | They propose to identify checkpoint models based on their respective capabilities and leverage them as data mixers by using their aggregated first-order influence approximation over source data. |
| Outcome: | The proposed framework shows significant improvements on eight reasoning benchmarks, with accuracy increases of up to 1.93%. |